skip to main content


Search for: All records

Creators/Authors contains: "Zhang, Jingyu"

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. In this work, we explore a useful but often neglected methodology for robustness analysis of text generation evaluation metrics: stress tests with synthetic data. Basically, we design and synthesize a wide range of potential errors and check whether they result in a commensurate drop in the metric scores. We examine a range of recently proposed evaluation metrics based on pretrained language models, for the tasks of open-ended generation, translation, and summarization. Our experiments reveal interesting insensitivities, biases, or even loopholes in existing metrics. For example, we find that BERTScore is confused by truncation errors in summarization, and MAUVE (built on top of GPT-2) is insensitive to errors at the beginning or middle of generations. Further, we investigate the reasons behind these blind spots and suggest practical workarounds for a more reliable evaluation of text generation. We have released our code and data at https://github.com/cloudygoose/blindspot_nlg. 
    more » « less
    Free, publicly-accessible full text available July 1, 2024
  2. Mononuclear non-heme iron enzymes are a large class of enzymes catalyzing a wide-range of reactions. In this work, we report that a non-heme iron enzyme in Methyloversatilis thermotolerans , OvoA Mtht, has two different activities, as a thiol oxygenase and a sulfoxide synthase. When cysteine is presented as the only substrate, OvoA Mtht is a thiol oxygenase. In the presence of both histidine and cysteine as substrates, OvoA Mtht catalyzes the oxidative coupling between histidine and cysteine (a sulfoxide synthase). Additionally, we demonstrate that both substrates and the active site iron's secondary coordination shell residues exert exquisite control over the dual activities of OvoA Mtht (sulfoxide synthase vs. thiol oxygenase activities). OvoA Mtht is an excellent system for future detailed mechanistic investigation on how metal ligands and secondary coordination shell residues fine-tune the iron-center electronic properties to achieve different reactivities. 
    more » « less
  3. Abstract Many measurements at the LHC require efficient identification of heavy-flavour jets, i.e. jets originating from bottom (b) or charm (c) quarks. An overview of the algorithms used to identify c jets is described and a novel method to calibrate them is presented. This new method adjusts the entire distributions of the outputs obtained when the algorithms are applied to jets of different flavours. It is based on an iterative approach exploiting three distinct control regions that are enriched with either b jets, c jets, or light-flavour and gluon jets. Results are presented in the form of correction factors evaluated using proton-proton collision data with an integrated luminosity of 41.5 fb -1 at  √s = 13 TeV, collected by the CMS experiment in 2017. The closure of the method is tested by applying the measured correction factors on simulated data sets and checking the agreement between the adjusted simulation and collision data. Furthermore, a validation is performed by testing the method on pseudodata, which emulate various mismodelling conditions. The calibrated results enable the use of the full distributions of heavy-flavour identification algorithm outputs, e.g. as inputs to machine-learning models. Thus, they are expected to increase the sensitivity of future physics analyses. 
    more » « less